Papers with explanation methods

15 papers
Explanation in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events.
Approach: They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations.
Outcome: The proposed methods are based on the models of large language models (LLMs) and their opaque nature.
Interpreting Language Models with Contrastive Explanations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding .
Approach: They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another.
Outcome: The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers .
Considering Likelihood in NLP Classification Explanations with Occlusion and Language Modeling (2020.acl-srw)

Copied to clipboard

Challenge: Existing explanation methods produce invalid or syntactically incorrect data, neglecting the improved abilities of recent NLP models.
Approach: They propose an explanation method that combines occlusion and language models to sample valid and syntactically correct replacements with high likelihood, given the context of the original input.
Outcome: The proposed method can sample valid and syntactically correct replacements with high likelihood, given the context of the original input.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement (P18-1)

Copied to clipboard

Challenge: a number of post hoc explanation methods for deep neural networks have been proposed . due to the complexity of the DNNs they explain, these methods are necessarily approximations and come with their own sources of error.
Approach: They propose two evaluation paradigms that cover two important classes of NLP problems . they propose LIMSSE, LRP and DeepLIFT as the most effective explanation methods .
Outcome: The proposed methods are most effective for explaining deep neural networks in NLP . the proposed methods can explain complex models without manual annotation .
Evaluating Explanation Methods for Neural Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Neural machine translation (NMT) has seen great success during recent years.
Approach: They propose a metric that measures the fidelity of explanation methods on translation tasks . they use an efficient approximation to evaluate several explanation methods .
Outcome: The proposed metric is efficient and can be used on translation tasks.
PromptExplainer: Explaining Language Models through Prompt-based Learning (2024.findings-eacl)

Copied to clipboard

Challenge: Existing explanation methods rely on linear approximations, accentuating irrelevant input tokens.
Approach: They propose a method that aligns the explanation process with the masked language modeling task of pretrained language models and leverages prompt-based learning to generate class-dependent explanations.
Outcome: Extensive experiments show that PromptExplainer outperforms state-of-the-art explanation methods.
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing explanation methods for image classification struggle to provide faithful and plausible explanations for predictions.
Approach: They propose a natural language explanation method that can be applied to any CNN-based classifier without altering its training process or affecting predictive performance.
Outcome: The proposed method can be applied to any CNN-based classifier without altering its training process or affecting predictive performance.
Lifelong Explainer for Lifelong Learners (2021.emnlp-main)

Copied to clipboard

Challenge: Existing explanation methods are inefficient when explaining a static black-box model.
Approach: They propose a Lifelong Explanation approach that continuously trains a student explainer under the supervision of a teacher on different tasks undertaken in LL.
Outcome: The proposed approach can be extended to include a teacher and maintain the same level of faithfulness to the black-box model as the student explainer while being up to 102 times faster at test time.
Alignment Rationale for Natural Language Inference (2021.acl-long)

Copied to clipboard

Challenge: Existing explanation methods pick prominent features, but alignments between words or phrases are more enlightening clues to explain the model.
Approach: They propose a method to generate alignment rationale explanations for co-attention based models in NLI by feature selection.
Outcome: The proposed method is more faithful and human-readable compared with existing methods.
Sequential Integrated Gradients: a simple but effective method for explaining language models (2023.findings-acl)

Copied to clipboard

Challenge: Existing explanation methods such as Integrated Gradients (IG) produce a path for each word of a sentence simultaneously, which can lead to sentences with no clear meaning or a significantly different meaning compared to the original one.
Approach: They propose to use a sequenced integrated gradient method to fix every word in a sentence and move it along a straight path to the word of interest.
Outcome: The proposed method improves on Integrated Gradients (IG) and DIG, but can produce sentences with different meanings than the original one.
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? (2020.acl-main)

Copied to clipboard

Challenge: a new study examines the impact of algorithmic explanations on simulatability of machine learning models . a model is simulatable when a person can predict its behavior on new inputs .
Approach: They conduct human subject tests to isolate effect of algorithmic explanations on simulatability . they find ratings of explanations are not predictive of how helpful they are .
Outcome: The results provide the first reliable estimates of how explanations influence simulatability . they show that ratings are not predictive of how helpful explanations are .
Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Training (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing explanation methods that generate keywords may be less effective due to missing critical contextual information.
Approach: They propose a new method to generate explanations for possible labels using LLMs and a dialectical prompt.
Outcome: The proposed method significantly improves accuracy and explanation quality over state-of-the-art methods on multiple datasets from diverse domains.
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown strong capabilities, enabling concise, context-aware answers in question answering tasks.
Approach: They propose a framework that provides valid uncertainty guarantees for LLMs . they also propose 'model-agnostic' uncertainty estimation method that maintains valid guarantees even under noise.
Outcome: The proposed method provides valid uncertainty guarantees even under noise.
Quantifying and Understanding Uncertainty in Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for estimating generation uncertainty do not provide finite-sample guarantees for reasoning-answer generation.
Approach: They propose a method that provides the uncertainty of the reasoning-answer structure with statistical guarantees.
Outcome: The proposed method disentangles reasoning quality from answer correctness while establishing theoretical guarantees for efficient explanation methods.
From Nodes to Narratives: Explaining Graph Neural Networks with LLMs and Graph Context (2026.acl-long)

Copied to clipboard

Challenge: Existing explanation methods for graph neural networks struggle to generate interpretable, fine-grained rationales.
Approach: They propose a lightweight framework that uses large language models to generate interpretable explanations for GNNs.
Outcome: The proposed framework generates interpretable explanations for GNN predictions using large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations